Papers with reasoning enhancement
MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances have enabled MLLMs to tackle complex challenges such as mathematical reasoning and multimodal understanding. |
| Approach: | They propose a multimodal refinement benchmark to evaluate the refinement capabilities of Multimodal Large Language Models (MLLMs) the benchmark categorizes errors into six error types to highlight areas for improvement in effective reasoning enhancement. |
| Outcome: | The proposed framework evaluates the refinement capabilities of multimodal large language models across six scenarios. |
Why Did Apple Fall: Evaluating Curiosity in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for evaluating curiosity-like behaviors in large language models lack curiosity-inspired features. |
| Approach: | They propose a psychology-inspired framework to evaluate curiosity in large language models . they adapt the Five-Dimensional Curiosity scale Revised (5DCR) to LLMs . |
| Outcome: | The proposed framework evaluates curiosity in large language models using questionnaires and behavioral studies. |